Repository navigation
README: DeepSWE benchmark results - #1511
Merged
Merged
Conversation
The Benchmarks section now carries the ten-harness DeepSWE comparison: the chart, a table of solved, cost per task, cost per solved issue and time, the limits of a one-seed run, and the V4.1 Flash and Kimi K3 runs. The per-harness numbers move into docs/benchmarks/deepswe. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE (cherry picked from commit d8f8d3d)
The README keeps the chart and one paragraph; the table, the later V4.1 Flash and Kimi K3 runs, the setup and the limits live in docs/benchmarks/deepswe. The chart's footer now says only "same model". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE (cherry picked from commit 9bdf25e)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE (cherry picked from commit 7c92173)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE (cherry picked from commit b1ac028)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE
The installer printed a three-line telemetry notice after the receipt. It prints none now; the binary's full notice still arrives before the first session's events are sent. The local install marker stays. The test asserts the installer prints no notice, and docs/TELEMETRY.md loses the installer's block. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE (cherry picked from commit 53e9fd9)
Nothing the installer prints mentions telemetry now; the marker is still written when it can be. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE (cherry picked from commit df57499)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE
AbirAbbas
added a commit
that referenced
this pull request
Sep 25, 2026
AbirAbbas
pushed a commit
that referenced
this pull request
Sep 25, 2026
…1436) A task's crew is picked per task by internal/crewroute (class, price on every connected route, quality minus lambda times cost) instead of a stored preset. /crew holds only what is allowed, what is pinned and a daily cap; --best and --cheap move one task. Old profiles migrate once. Route health, a per-seat fallback ladder, per-task and daily spend limits, /redo stronger, remote protocol 18. Added in review: a pin keeps its thinking level, auto in the reflex or small-work row reads its default and migrates, -yes-spend passes the daily cap, the cap refusal no longer offers --cheap, zero crew money is not drawn, and the manual describes the one seats row; dev merged in over #1429, #1485, #1494 and #1511. Co-authored-by: Santosh kumar <29346072+santoshkumarradha@users.noreply.github.com> Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AbirAbbas
added a commit
that referenced
this pull request
Sep 25, 2026
…side senior-dev Brings in #1494 (team delegation), #1511 and #1436 (per-task worker, planner and checker routing). A senior-dev run keeps its own road beside them: it works alone in the folder it holds, so a via proposal starts one program run and does not go through the team split; it keeps the models it was asked for, or the pinned or profile worker, and #1436's per-task routing and $5 task cap apply to codeaf's own tasks, while senior-dev keeps its finite per-run ceiling. Both sides had claimed remote protocol version 18; the combined wire is version 19. Program badges fit #1494's grouped side column, the manuals describe the merged model choice, and the prompt-size ledger records the combined fixed prefix (57,218 bytes). dev's unbounded-launch test now holds its six workers at a barrier instead of guessing with a sleep, which failed on a loaded machine. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AbirAbbas
pushed a commit
that referenced
this pull request
Sep 28, 2026
…1635) The installer on main, on top of v0.4.1 and nothing else from dev: checked steps with a spinner, a link into a folder already on PATH so codeaf works in the same terminal, a Get started guide, and on a terminal the question `Start codeaf in <folder> now? [Y/n]`, read from /dev/tty. The binary is untouched; only the script, its tests and the pages that describe it move. scripts/install.sh, test/installer-telemetry.sh and internal/release/install_test.go are byte-identical to #1635's head, which also carries the installer halves of two dev changes it was written on: the installer prints no telemetry notice (#1511, with the matching docs/TELEMETRY.md paragraph; v0.4.1 still prints the full notice itself before any count is sent), and an install under another name puts that name first in its receipt (#1519). docs/GUIDE.md and the two manual pages take #1635's paragraphs in place of the ones that described the three-line notice. It also carries the review fixes made on the pull request before it landed: only the installer's own link on PATH is ever replaced, the start question is not asked in the home folder or at /, the spinner's frames survive a UTF-8 locale that is not installed, and the guide sends DeepSeek, Qwen, GLM, Kimi and Ollama through /connect, since v0.4.1's first run offers OpenRouter alone. This commit is merged into staging and dev so main stays an ancestor of staging and staging of dev, and every later promotion is still a fast-forward. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AbirAbbas
added a commit
that referenced
this pull request
Sep 28, 2026
Brings e446507 (the guided installer, on top of v0.4.1) into staging, so main can fast-forward onto it and the rc it publishes passes the "already on staging" check. Staging already carried #1511's half of that commit; the conflicts in scripts/install.sh and test/installer-telemetry.sh resolve to #1635's head, byte for byte. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
#benchmarks./senior-dev… First of ten harnesses on DeepSWE".assets/readme/benchmark-deepswe.webp), one paragraph, and a link todocs/benchmarks/deepswe/.docs/benchmarks/deepswe/: the ten-harness table, the V4.1 Flash and Kimi K3 runs, setup and limits (README.md), and one row per harness (arms.csv), from the DeepSWE comparison run of 2026-09-12.test/installer-telemetry.shpasses, 28 of 28.Cherry-picked from
zeropoint95/senior-dev-harness(#1488) without its code.🤖 Generated with Claude Code
https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE